Papers by Alexander M Rush
Challenges in Trustworthy Human Evaluation of Chatbots (2025.findings-naacl)
Copied to clipboard
| Challenge: | apathetic or adversarial annotators can corrupt the reliability of open leaderboard rankings . human annotation is widely accepted as the gold standard for open-ended text generation tasks . |
| Approach: | They show that bad annotations can corrupt the reliability of open leaderboard rankings . they argue that human annotation is widely accepted as the gold standard . |
| Outcome: | The proposed algorithm can corrupt the reliability of open leaderboard rankings by up to 5 places. |
Great Memory, Shallow Reasoning: Limits of kNN-LMs (2025.naacl-short)
Copied to clipboard
| Challenge: | Existing models trained on poor quality data have shown strong performance in language modeling and some downstream benchmarks. |
| Approach: | They evaluate kNN-LMs on a diverse set of tasks and evaluate their performance. |
| Outcome: | The proposed extension could improve on a variety of tasks, but it fails to perform on reasoning tasks that require integrating multiple pieces of information. |